Skip to content

feat(recipe): add Qwen3.5 text MXFP8 long-context SFT - #6300

Merged
cuichenx merged 5 commits into
NVIDIA-NeMo:mainfrom
cuichenx:chcui/maya/qwen35-text-mxfp8-sft
Oct 7, 2026
Merged

cuichenx merged 5 commits into
NVIDIA-NeMo:mainfrom
cuichenx:chcui/maya/qwen35-text-mxfp8-sft

Conversation

@cuichenx

@cuichenx cuichenx commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Adds a Qwen3.5-35B-A3B text-only 128K MXFP8 SFT recipe for 16 GB200 GPUs, with recipe exports and focused tests. The recipe uses TP1/CP8/EP16, one MTP layer, packed CoderForge data, MXFP8 parameter gather and gradient-buffer reuse, CuTeDSL grouped MLP, single grouped weights, and HybridEP. Model and dataset revisions are pinned.

The recipe starts directly from _sft_common() and retains workload-specific settings. Redundant assignments now inherit the common builder, model provider, dataset, and precision defaults.

Validation: 105 focused text/Qwen recipe tests and all-file pre-commit pass on this refactor. Before/after normalized configurations match with offline HF metadata and the real model provider. Independent subagent review found no blocking issues. This PR contains only the new recipe, exports, and focused tests; no verification card or documentation changes.

Historical launcher-versus-Bridge comparisons completed 100 updates with matching losses, but used an uncommitted recipe precursor and Megatron-LM PR #7611. They are not clean-checkout verification of this PR. Masked MTP correctness requires NVIDIA/Megatron-LM#7611 or an equivalent fix in the Bridge MCore pin; this PR does not update dependencies.

Signed-off-by: Chen Cui <chcui@nvidia.com>
@cuichenx cuichenx added feature New capabilities, enhancements, or enablement work area:recipe Training recipes and launch configs needs-review PR is ready for code review and waiting on a reviewer labels Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Automatic Claude reviews have been retired. To request a pull-request review, post a comment containing:

/review

Add model=claude to use a Claude reviewer (the default is model=codex). mode=light|strict selects the review depth; for example, /review model=claude mode=strict. Comment /review help for all options.

@cuichenx cuichenx added blocked Work cannot move forward until an external dependency is cleared and removed needs-review PR is ready for code review and waiting on a reviewer labels Oct 2, 2026
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
Signed-off-by: Chen Cui <chcui@nvidia.com>
@cuichenx
cuichenx merged commit 7bbe5de into NVIDIA-NeMo:main Oct 7, 2026
94 checks passed
ilml added a commit to ilml/Megatron-Bridge that referenced this pull request Oct 8, 2026
Resolve conflicts with the Qwen3.5 text MXFP8 long-context SFT recipe from
main (NVIDIA-NeMo#6300). Both sides registered a new recipe in recipes/qwen/__init__.py
and recipes/qwen/gb200/__init__.py, so both exports are kept. The qwen35.py
import block takes main's imports, which include everything this branch needs.

Signed-off-by: Tom Long <tolong@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
test — 9952be17 Deployed Oct 7, 2026 by copy-pr-bot[bot] via cicd-wait-in-queue #22562
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:recipe Training recipes and launch configs blocked Work cannot move forward until an external dependency is cleared feature New capabilities, enhancements, or enablement work

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants